OEIS Open:语言模型能把多少猜想变成定理?
文章背景与核心概要
本文介绍了 OEIS Open,这是一个旨在评估语言模型(LMs)自主解决开放性数学猜想能力的新基准。该基准包含来自“整数数列在线百科全书”(OEIS)的 492 个开放性猜想,并在 Lean 中进行了形式化,为通用语言模型提供了一个开源、安全的测试环境。
核心研究结果表明,配备极简工具并在适度预算(每次尝试 50 美元)下运行的语言模型,成功解决了 30%(147 个)的猜想。此外,在名为 OEIS Open Lite 的子集上对表现最好的语言模型进行评估,将每次尝试的预算提高到 200 美元时,成功率达到了 44%。有趣的是,为语言模型增加 476,000 篇 arXiv 数学论文或采用更复杂的智能体循环并没有提升性能。尽管这些特定的猜想在数学上的重要性各有不同,但结果证明语言模型能够以低成本开展自主的开放性研究问题。
执行摘要
Executive Summary
This paper introduces OEIS Open, a new benchmark designed to evaluate the ability of language models (LMs) to autonomously resolve open mathematical conjectures. Comprising 492 open conjectures from the On-Line Encyclopedia of Integer Sequences (OEIS)—formalized in Lean—the benchmark provides an open-source, secure testing environment for generic LMs.
Key findings reveal that LMs equipped with minimal tools and operating on a modest budget ($50 per attempt) successfully resolved 30% (147) of the conjectures. Furthermore, evaluating the top-performing LM on a subset called OEIS Open Lite at an increased budget of $200 per attempt yielded a 44% success rate. Interestingly, augmenting LMs with 476,000 arXiv mathematics papers or employing more sophisticated agent loops did not improve performance. While these specific conjectures vary in mathematical significance, the results demonstrate that LMs can tackle autonomous open research problems at low costs.
论文元数据
Paper Metadata
- arXiv 标识符:
arXiv:2608.11941[cs.AI]- 作者: Tom Adamczewski
- 提交时间: 2026年8月12日
- 主要学科: 人工智能 (
cs.AI) - MSC 分类: 68V15 (主要), 68T07
- ACM 分类: I.2.3; F.4.1
核心亮点与发现
Key Highlights & Findings
- 基准构建: 基于 OEIS 中的 492 个开放数学猜想构建,由 Tsoukalas 等人在 Lean 定理证明器中进行了形式化。
- 成本效益: 使用基础工具集的语言模型在每次尝试 50 美元预算 内成功解决了 147 个猜想。
- OEIS Open Lite: 对 100 个随机猜想组成的较小子集进行的评估表明,在 200 美元预算 下,当前表现最好的语言模型取得了 44% 的成功率。
- 外部文献影响: 让语言模型访问近 50 万篇来自 arXiv 的研究论文并没有提升性能,复杂的智能体循环同样没有带来提升。
访问与资源
Access & Resources
- 全文选项:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 代码与结果仓库:
- GitHub 评估代码
- GitHub 运行结果
- 许可证: 知识共享署名 4.0
